BMC Medical Research Methodology
○ Springer Science and Business Media LLC
Preprints posted in the last 90 days, ranked by how well they match BMC Medical Research Methodology's content profile, based on 47 papers previously published here. The average preprint has a 0.06% match score for this journal, so anything above that is already an above-average fit.
de Carvalho, F. R.; Gavaia, P. J.
Show abstract
Purpose The application of machine learning (ML) to osteoporosis prediction has expanded rapidly, yet no comprehensive meta-analysis has synthesized the discriminative performance of these models across all ML categories, data types, and validation strategies. This systematic review and meta-analysis aimed to evaluate the diagnostic and predictive accuracy of ML and deep learning models for osteoporosis prediction in adult populations. Methods Systematic searches of PubMed, Embase, Web of Science, and IEEE Xplore were conducted for studies published between January 2020 and February 2026. Studies developing, validating, or applying ML models for predicting osteoporosis, low bone mineral density, or osteoporotic fractures in adults were included. Methodological quality was assessed using the Prediction Model Risk of Bias Assessment Tool (PROBAST). Area under the receiver operating characteristic curve (AUC) values were pooled using random-effects meta-analysis with logit transformation. Subgroup analyses were performed by data type, ML category, external validation status, and population type. The review followed PRISMA 2020 guidelines. Results Thirty-three studies were included in the qualitative synthesis and 27 in the meta-analysis. The pooled AUC was 0.879 (95% CI: 0.853 0.901), with substantial heterogeneity (I = 99.5%). Imaging-based models outperformed clinical data models (AUC = 0.905 vs. 0.872). Deep learning achieved the highest pooled AUC (0.909), followed by ensemble methods (0.874) and traditional ML (0.840). Externally validated models showed lower performance than internally validated ones (AUC = 0.868 vs. 0.897). PROBAST assessment rated 32 of 33 studies (97.0%) as low risk of bias, though this proportion should be interpreted cautiously given that PROBAST was designed for traditional prediction models and may not fully capture ML-specific sources of bias. Egger's test indicated significant publication bias (p < 0.001). Explainable AI methods were employed in 60.6% of studies, identifying age, body weight, and alkaline phosphatase as the most frequent top predictive features. Conclusions Machine learning models demonstrate overall good discriminative performance for osteoporosis prediction, albeit with substantial heterogeneity across studies (I = 99.5%), and show potential as complementary screening tools, particularly in settings with limited DXA access. Deep learning models applied to imaging data and ensemble methods using clinical variables achieved the strongest subgroup estimates. However, extreme heterogeneity, evidence of publication bias, and limited prospective validation warrant cautious interpretation of the pooled estimate. Future research should prioritise multi-centre external validation, standardised reporting following TRIPOD+AI guidelines, and prospective clinical trials to establish real-world clinical impact.
Di Carluccio, E.; Koliopanos, G.; Ojeda, F. M.; Weimar, C.; Ziegler, A.
Show abstract
Statistical prediction models for binary outcomes are becoming increasingly popular. One significant challenge is calibrating these models to suit the characteristics of a target population that is structurally different from the original population. Calibration is especially challenging when there is no training data available from the target population. To address this problem, we propose a novel calibration method, SimCal, which uses synthetic data generated from the model development data in conjunction with marginal statistics from the calibration cohort. We show that expert judgment modeling (EJM) may be used for calibration if cross-sectional data from the target population are available comprising expert judgments about the potential outcome and the covariates. We describe three alternative calibration approaches when calibration data are lacking: similarity-binning averaging (SBA), adaptive calibration of predictions (ACP), and Elkan calibration. In a simulation study, we compare SBA, ACP, Elkan calibration, and SimCal. R code for applying these methods is provided from the re-analysis of data on coronary artery disease. We illustrate all 5 calibration approaches with a real data set for predicting functional outcome after stroke and all approaches but EJM in the re-analysis of the Cleveland Clinic data. None of the approaches performed convincingly well in all situations. SimCal performed well when model parameters were correctly specified. EJM failed on the stroke data. Further research is urgently required for calibration in the absence of calibration data.
Wang, Z.; liu, y.
Show abstract
Background Primary care in China lacks structured mental-health assessment, and the machine-learning models that could support such screening are typically developed on heavily selected samples. Cumulative inclusion and exclusion criteria, though usually treated as neutral data-cleaning steps, can create heterogeneity in predictive reliability among retained participants. Using the China Health and Retirement Longitudinal Study (CHARLS) 2011 baseline, we quantified how selection funnels distort epidemiological associations and inflate machine-learning metrics, and tested selective prediction as mitigation. Methods Using the CHARLS 2011 baseline with temporal external validation in CHARLS-2018, we built a four-level selection funnel (L0-L3), evaluated five classifiers with nested cross-validation and SMOTE, and compared model-embedded uncertainty with a decoupled predictor-selector framework; XGBoost cross-validation residuals drove risk stratification and classification and regression tree (CART) rules. Results Sample sizes fell from L0 n=17,705 to L3 n=4,256 (24.0%). The cancer-depression odds ratio attenuated from 1.78 (95% CI 1.32-2.41) to 1.39 (0.74-2.63), losing significance. AUC rose with selection but not after multiple-comparison correction, whereas calibration error increased for four of five models. Model-embedded uncertainty succeeded only for XGBoost; with the decoupled XGBoost residual selector, all five models achieved selective prediction at approximately 20% coverage (test AUC 0.90, 95% CI 0.85-0.95), abstaining on approximately 80% of cases for individual safety. Risk stratification was stable (residual Spearman correlations >0.95; multi-seed Jaccard 0.88), and CART rules used self-rated health, education, pain, and marital status. Conclusions The findings support a deployable primary-care triage pathway: a four-variable rule identifies patients suitable for algorithm-assisted scoring (approximately 20% coverage) and routes the remainder to human evaluation. Methodologically, cumulative selection bias produces a dual distortion: epidemiological associations are compressed and machine-learning metrics inflated. Selective prediction is limited mainly by uncertainty-indicator design. Performance metrics should be reported with selection level, coverage, and calibration trajectory. Decoupled selective prediction with CART rule extraction provides an actionable framework for quality-controlled, tiered-care deployment. Keywords: selective prediction, selection bias, CHARLS, depression, predictor-selector decoupling, uncertainty quantification, classification and regression tree, triage, clinical decision support, health management.
Hendrickx, N.; Mentre, F.; Karlsson, M. O.; Hooker, A. C.; Traschütz, A.; Schüle, R.; PROSPAX Consortium, ; EVIDENCE-RND Consortium, ; Synofzik, M.; Comets, E.
Show abstract
We propose two new tests to detect drug effects (DE) in trials of one to very few patients followed during two periods (before and after initiation of a treatment). Both methods use longitudinal natural history data to inform the estimation of each patient's DE. The first method uses a non linear mixed effect model (NLMEM) reflecting an expected natural history with a hypothetical drug effect, to estimate the Conditional Distribution of the Drug Effect (CDDE). The second method trains a Pareto Depth Analysis (PDA) algorithm, a machine learning based approach based on outlier detection, that we implement using data simulated under the NLMEM. We evaluated the two tests with a simulation study. We used data from the PROSPAX study in Autosomal Recessive Cerebellar Ataxias (ARCAs, to derive a NLMEM for the Scale for the Assessment and Rating of Ataxia score. The CDDE method provided controlled type I error and, in some scenarios, adequate corrected power, though sensitivity analyses showed vulnerability to misspecification. The PDA method demonstrated lower statistical power except with high score precision. These results highlight different strategies for quantifying treatment effects in ultra rare, patient' specific trials. They can inform methodological design for future ARCA precision therapies.
Lytras, T.; Athanasiadou, M.
Show abstract
Background: Reliable estimation of excess mortality is central to population health surveillance. We introduce NeMMo (New Mortality Model), an evolution of the EuroMOMO model for estimating weekly all-cause expected mortality, and assess its behaviour and performance on empirical data. Methods: NeMMo incorporates population offsets, stratifies observed deaths by age group and models seasonality using a periodic B-spline rather than a Serfling-type sinusoidal function. Baseline weeks are selected by a data-driven procedure minimizing the skewness of the residuals before refitting the model, instead of relying solely on fixed calendar windows. NeMMo enables pooling across age groups, direct age standardization and incorporation of external predictors. We applied NeMMo and EuroMOMO to mortality and population data downloaded from Eurostat for 31 countries from 2015 onwards, excluding the COVID-19 pandemic period from baseline estimation. Results: For most countries NeMMo produced a higher expected mortality baseline that better tracked observed deaths, as well as tighter prediction intervals and higher maximum Z-scores, suggesting improved discrimination of mortality excesses. Z-scores and P-scores during non-pandemic weeks were closer to zero with NeMMo than with EuroMOMO but further elevated during pandemic weeks, providing greater separation between pandemic and non-pandemic mortality. Incorporating population offsets resulted in negative linear trends across all countries, consistent with declining mortality after accounting for demographic changes. The periodic B-spline identified substantial heterogeneity in the shape and timing of seasonal mortality that was not captured by a sinusoidal function. Conclusions: NeMMo provides a flexible and parsimonious framework for all-cause mortality surveillance that improves the established EuroMOMO model and offers theoretical, empirical and practical advantages. It is thus suitable both for detecting short-term spikes and for the long-term, age-adjusted quantification and comparison of mortality excesses that has become increasingly important since the COVID-19 pandemic. The accompanying 'nemmo' package for R facilitates its widespread adoption and application.
Elson, R.; McIntyre, K. M.; Hardingham, M. B.; Luechtefeld, T.; Lake, I. R.
Show abstract
Abstract Climate change is altering environmental conditions that influence foodborne disease transmission, yet traditional systematic reviews cannot keep pace with expanding evidence. We assessed whether an LLM-assisted workflow could generate a rapid, repeatable, and policy-relevant living evidence base for climate-sensitive foodborne disease. We combined structured PubMed searches (2010-2023), gold-standard human labelling, and iterative refinement of a GPT?4?Turbo?based auto-labeller within the SysRev platform. Pathogens of public-health importance in England were selected a priori. Model performance was evaluated against human reviewers using recall, precision, specificity, accuracy, and balanced accuracy. The refined inclusion model achieved 89{middle dot}2% recall, 59{middle dot}2% precision, 84{middle dot}5% specificity, and 85{middle dot}4% accuracy across 1,044 screened abstracts, identifying 436 studies for inclusion. Post-hoc re-evaluation of discordant abstracts showed that records excluded by the model but included during initial human screening did not meet the refined inclusion criteria. Frequently identified climate exposures included rainfall, temperature, seasonality, and humidity; norovirus, Salmonella, Campylobacter, and Cryptosporidium were the most common pathogens. An LLM-assisted workflow can generate living evidence for climate-sensitive foodborne disease with high recall and improved screening consistency. The approach is scalable, auditable, and suitable for secure institutional environments, supporting horizon scanning and climate-health risk assessment.
Katarynczuk, K.; Stachowiak, A.; Piorkowska, N. J.; Ostromecki, A.; Franik, G.; Bizon, A.
Show abstract
Background: Machine-learning models for polycystic ovary syndrome (PCOS) and other conditions frequently report near-perfect diagnostic performance, but retrospective datasets assembled from routine clinical practice can encode diagnostic-group membership in how data were acquired rather than in disease biology, and this acquisition-related information can be indistinguishable from genuine clinical signal under conventional validation. Objective: To determine, using a real-world PCOS cohort as a case study, whether high classification performance reflected clinically meaningful information or artifacts of data provenance, schema structure, and measurement-acquisition workflow, and to develop a generalizable audit framework for detecting such artifacts in retrospective medical machine learning. Methods: We analyzed 1,331 retrospective records (1,286 PCOS, 45 controls) from a single endocrine-gynecology database. A layered acquisition-bias framework compared classification performance using (i) raw and harmonized missingness patterns alone, (ii) measured values with and without explicit missingness indicators, and (iii) ascertainment-balanced feature sets with and without age. Logistic regression and random forest were evaluated using repeated stratified cross-validation, bootstrap resampling, label-permutation testing, and calibration analysis, and the framework was validated against a semi-synthetic experiment with known ground truth. Results: Diagnostic status was perfectly predicted (ROC-AUC = 1.000) from missingness patterns alone, before any clinical value was examined, and this persisted after semantic harmonization of duplicated source columns. Performance declined progressively as acquisition-sensitive information was removed, from near-ceiling in raw and harmonized value models to a mean ROC-AUC of approximately 0.80-0.82 in the most restrictive ascertainment-balanced, age-excluded representation. The semi-synthetic experiment reproduced this pattern under known data-generating conditions, confirming that harmonization removes schema-fragmentation artifacts but not workflow-driven acquisition bias. Conclusions: Apparent diagnostic performance in this cohort was substantially attributable to diagnostic workflow and data-acquisition structure rather than to a stable, transportable biological signal. The layered audit framework generalizes beyond PCOS and offers a practical tool for detecting acquisition-related leakage in retrospective clinical machine-learning studies.
LIn, H.; Lyu, J.
Show abstract
BackgroundQuality Control Circle (QCC) reports are often reviewed qualitatively, but reviewer workload and inter-rater variability make large-scale assessment difficult. We evaluated whether multiple large language models (LLMs) could score QCC methodological quality reliably on a designed-anchor benchmark. ObjectiveTo estimate inter-model reliability for QCC quality scoring and to assess whether model scores align with designed synthetic anchors and remain descriptively comparable to a small set of public PMC QCC reports. MethodsWe evaluated 30 synthetic QCC reports and 8 public PMC QCC reports across four primary evaluators (GPT, Gemini, Grok, DeepSeek) and one sensitivity evaluator (Claude); Claude was excluded from the primary panel because it shared the model family used during prompt development. Each synthetic case was scored across eight QCC quality dimensions in three runs per evaluator. We summarized each evaluator by median scores, then estimated ICC(A,1) across the primary panel. We also examined score-based calibration against designed anchors, keyword-assisted defect mention, leave-one-out and k=5 sensitivity, and a descriptive synthetic-versus-PMC distributional plausibility check. ResultsInter-model reliability on the primary k=4 panel was excellent: ICC(A,1) = 0.953 (95% CI 0.944 to 0.962) with 237 pooled case-dimension rows. The pre-specified k=5 sensitivity analysis including Claude was 0.954, and leave-one-out estimates within the primary panel ranged from 0.950 to 0.959. Score-based calibration against designed anchors met the prespecified target in 57/58 trap-affected case-dimension rows (98.3%). Keyword-assisted defect mention was present in 51/58 trap instances (87.9%). The synthetic-versus-PMC comparison was descriptively similar across all eight dimensions, and all dimensions met the predefined descriptive margin check. ConclusionsIn this designed-anchor pilot, multi-model LLM scoring of QCC methodological quality showed high inter-model reliability and stable alignment with synthetic anchor scores. These findings support benchmark feasibility, but they do not establish expert validity, clinical validity, or operational deployment readiness.
Piorkowska, N. J.; Ostromecki, A.; Franik, G.; Bizon, A.
Show abstract
Background Unsupervised machine learning has become a cornerstone of computational phenotyping across clinical medicine, genomics, imaging, and multi-omics research. However, phenotype discovery relies on a sequence of analytical decisions - including missing-data handling, preprocessing, dimensionality reduction, clustering methodology, and stochastic initialization - that are rarely evaluated collectively. Although clustering stability has been extensively investigated, the robustness of complete analytical workflows remains largely unexplored. Results We developed an Analytical Perturbation Framework that systematically quantifies the robustness of phenotype discovery by perturbing complete unsupervised learning workflows rather than individual clustering algorithms. Using a real-world cohort of 1,286 women with polycystic ovary syndrome (PCOS), we generated 116 valid analytical pipelines comprising alternative preprocessing strategies, missing-data handling methods, dimensionality reduction approaches, clustering algorithms, and random initializations. Agreement between independently generated phenotype solutions was consistently low (median Adjusted Rand Index = 0.079), indicating substantial sensitivity of phenotype discovery to routine analytical decisions. Variance decomposition identified preprocessing as the largest contributor to phenotype instability (22.8%), followed by clustering methodology (14.6%), whereas stochastic initialization explained only 3.1% of the observed variability. At the patient level, most individuals exhibited reproducible phenotype assignments (median Patient Robustness Score = 0.719), although a substantial subgroup showed markedly lower assignment stability. Feature perturbation analyses identified follicle-stimulating hormone, anti-thyroglobulin antibodies, anti-thyroid peroxidase antibodies, total testosterone, luteinizing hormone, and androstenedione as the strongest contributors to computational robustness, rather than biological importance. Finally, phenotype solutions demonstrating greater computational robustness also exhibited greater biological coherence during independent validation.
Savu, A.; Dover, D. C.; Hajihosseini, M.; Gaudet, L. A.; Kaul, P.
Show abstract
Background and Objective. Missing data frequently occurs in health databases and can bias analyses if not correctly dealt with. Using real-world data, we compared complete-case and multiple-imputation methods for recovering true parameters of a multivariable logistic regression model for the association between maternal glucose levels during pregnancy and child excess weight at preschool age, where missing values were present in as much as 30% of our sample. Methods. This study utilized a cohort of 130,424 children with complete preschool-age body mass index (BMI) measurements from the Calgary and Edmonton health regions of Alberta, Canada. In the complete BMI data, we introduced missingness through deletion following three distinct mechanisms: missing completely at random (MCAR), at random (MAR), and not at random (MNAR). To handle the missing data created, we employed complete-case and multiple-imputation methods. Maternal glucose levels during pregnancy were categorized into five groups and its association with child excess weight at pre-school age was determined based on a logistic regression model using the full observed data (yielding true values), observed data that was not deleted (complete-case estimates), and imputed data (multiple-imputation estimates). The accuracy of complete-case and multiple-imputation estimates were evaluated against the true values. Finally, we conducted a sensitivity analysis for the MNAR mechanism using pattern-mixture models with an additive shift. Results. Under MCAR and MAR, multiple-imputation generally outperformed complete-case, yielding smaller absolute and relative bias. Both methods achieved high significance ([≥] 0.96) for most effects. Mean squared errors for multiple-imputation and complete-case were similar missing completely at random, missing at random, and coverage was consistently high ([≥] 0.99). Under MNAR, both complete-case and multiple-imputation showed poor performance regarding bias and statistical significance. Sensitivity analysis using pattern-mixture models indicated performance varied by specific effect. Conclusions. Under MCAR and MAR, multiple-imputation introduced higher bias but demonstrated superior overall performance based on mean squared error and restored statistical power. Conversely, both methods failed under MNAR, where pattern-mixture modeling sensitivity analyses revealed highly variable, effect-specific performance due to unverifiable shift assumptions. When faced with missing data, researchers should assess missingness mechanisms, report both complete-case and multiple-imputation estimates under MCAR/MAR while accounting for power-versus-bias tradeoffs, and employ pattern-mixture sensitivity analyses to test robustness when MNAR is plausible.
MA, Z.; XIANG, Y.; So, H.-C.
Show abstract
Abstract Purpose This study introduces a novel approach to address unmeasured confounding in terminal event studies using the prior event rate ratio (PERR) method. The proposed approach PERR_{proxy} used a proxy event to replace the original terminal event in the pre-exposure period, enabling the application of PERR in terminal event settings. Additionally, we also applied difference in difference (DID) regression, which is conceptually analogous to PERR to estimate the standard errors and confidence intervals of PERR_{proxy}. Methods We conducted numeric simulations to evaluate the validity of PERR_{proxy} approach and assessed its performance under varying levels of unmeasured confounding effects, baseline hazard ratios, and the correlation between the proxy and terminal events. To demonstrate its practical applicability, we also performed an empirical analysis to investigate the impact of severe hospitalized COVID-19 on circulatory system disease mortality using the PERR_{proxy}. Results In simulation studies, PERR_{proxy} effectively reduced the unmeasured confounding effects compared to the conventional methods. The performance of PERR_{proxy} was influenced by the strength of unmeasured confounding, baseline hazard ratios, and the correlation between the proxy and terminal outcomes. In addition, difference in difference (DID) regression had much faster computational speed for estimating standard errors and confidence intervals compared to bootstrap. In the empirical analysis, PERR_{proxy} identified that severe hospitalized COVID-19 as a significant risk factor for the circulatory system disease mortality and reduced the unmeasured confounding effects. Conclusions The PERR_{proxy} approach extends the applicability of the original PERR method to terminal event studies, offering a promising solution for addressing unmeasured confounding. Additionally, the DID regression framework provides a computationally efficient alternative for parameter estimation in PERR-based studies. However, careful consideration is still required in PERR_{proxy} for proxy events selection and other underlying assumptions of the PERR method to ensure valid results. Keywords: prior event rate ratio, unmeasured confounding, proxy event, terminal event study, observational study, electronic health records
Yahaya, Y.; Khan, S.; Rani Saha, P.; Meia, M. A. A.
Show abstract
Diagnosed diabetes affects approximately 38.4 million Americans, but its burden is not evenly distributed across U.S. counties. Existing machine-learning studies have mainly focused on individual risk prediction using biometric, clinical, or survey variables. These approaches are less suited to explaining why diagnosed diabetes prevalence differs geographically across counties. We developed an explainable gradient-boosting framework for predicting county-level diagnosed diabetes prevalence across 2,957 U.S. counties using an ecological cross-sectional design. The analysis integrated food-environment, socioeconomic, occupational, demographic, health-behavior, and clinical indicators from five public data sources. Four regression models were compared: Elastic Net, Random Forest, XGBoost, and LightGBM. LightGBM was selected as the primary model based on validation-set RMSE and interpreted using SHAP TreeExplainer. The validation-selected LightGBM model achieved a held-out test RMSE of 0.423 percentage points, R{superscript 2} = 0.964, and MAPE = 2.76%. Although XGBoost achieved a lower test RMSE of 0.399 and R{superscript 2} = 0.968, it was retained as a secondary benchmark because primary-model selection was based only on validation performance. A sensitivity model using only structural and contextual predictors, and excluding CDC PLACES health-behavior and clinical covariates, retained substantial predictive performance (R{superscript 2} = 0.827). Poverty rate was the most frequent dominant positive structural SHAP contributor nationally (n = 772 counties, 26.1%), followed by food insecurity rate (n = 707, 23.9%), Supplemental Nutrition Assistance Program (SNAP) participation rate (n = 316, 10.7%), unemployment rate (n = 224, 7.6%), and median household income (n = 178, 6.0%). Residual Morans I decreased from 0.665 to 0.069 after model fitting. Explainable machine learning using public county-level data can characterize geographic variation in diagnosed diabetes prevalence. County-level SHAP maps may support local hypothesis generation, but should be interpreted as explanations of model predictions rather than causal effects.
Wilner, L.; Casey, J. A.; Mooney, S. J.; Do, V.; Ma, Y.; Benmarhnia, T.; Dey, A. K.
Show abstract
When randomized controlled trials are infeasible, researchers may leverage natural experiments for causal inference. Interrupted time-series (ITS) designs compare observed post-event trends to counterfactual predictions from pre-event data. Two-stage ITS designs use flexible models to generate optimized counterfactual predictions in the first stage, then estimate intervention effects by comparing observed to predicted outcomes in the second stage. Fitting high-dimensional versions of these models is challenging, requiring systematic infrastructure to ensure rigor and reproducibility. In response, we developed its2s, an open-source Python package implementing the two-stage ITS design with machine learning. its2s allows users to specify an intervention date and training/testing periods, select among built-in model architectures (e.g., Prophet-XGBoost, NeuralProphet), and generate confidence intervals via moving block bootstrap, preserving temporal autocorrelation in residuals. its2s layers defaults, configuration files, and runtime overrides to support workflows ranging from rapid default implementations to highly tailored analyses. We validated its2s using two case studies: a simulation with a 12% policy effect, recovering the true effect as 11.77%, and an analysis of the 2021 Pacific Northwest heat dome, finding 53% excess injury mortality over the following three weeks. its2s provides a flexible, reproducible framework for ITS-based quasi-experimental research, lowering barriers to rigorous machine learning-based counterfactual modeling.
McLean, K. W.; LaBonte, J.; Macaulay, K.; Kassam-Adams, S.
Show abstract
This study documents the derivation and validation of a deterministic algorithm for cause-of-death (COD) ascertainment from longitudinal real-world medical claims data, evaluated against an independent state-level death certificate file. Death certificates are the dominant reference standard in mortality research but carry well-documented limitations, including primary-cause error rates estimated at 20-40\% across empirical studies. A matched analytic cohort of 216,382 individuals (Connecticut death records, 2017--2025, age 25 and above) was constructed after exclusion of mechanism-of-injury cases and removal of ill-defined symptom-code entries from both sources. Concordance between algorithmic and certificate-based COD was assessed through three complementary frameworks: age-stratified positive predictive value (PPV) at the ICD-10-CM chapter level under a full-set concordance scenario; mean absolute rank difference (MARD) for chapters identified by both sources; and analyses of breadth, depth, and code-level specificity of COD reporting. Chapter-level PPV was strongest for individuals aged 55 and above, with all estimates representing conservative lower bounds given the known error rate of the certificate reference standard. The algorithm consistently reported broader and more granular contributing cause profiles than the death certificate, with discordances directionally consistent with the well-documented tendency of certificates to under-report contributing conditions. These findings support the conclusion that algorithmic COD ascertainment from longitudinal claims data is a feasible and scalable alternative to certificate-based attribution and, at population scale, a principled methodology for characterising death certificate error rates beyond what small-sample chart review studies can achieve.
Xiang, S.; He, H.; Xie, Z.; Cheng, C.-Y.; Li, H.; Liu, D.
Show abstract
Agentic workflows can coordinate modelling, but balancing predictive performance, measurement burden and reproducibility is unclear. We developed DXA Agent, an agentic workflow for dual-energy X-ray absorptiometry (DXA) outcomes integrating planning, feature-model refinement, tools, provenance and hypothesis-generating interpretation. Models were independently developed and tested in UK Biobank (5,318 participants) and the National Health and Nutrition Examination Survey (NHANES; 3,777 participants), using cost-efficient and no-limit strategies. Across 20 UK Biobank and three NHANES bone mineral density sites, cost-efficient models achieved lower RMSE and higher R2 than the best conventional comparator, with median relative RMSE reductions of 10.9% and 9.9%, respectively. Classification was task dependent: UK Biobank osteoporosis averaged AUROC 0.839 and PR-AUC 0.182, whereas NHANES performance was comparable with conventional models. Higher-burden features did not consistently improve prediction. These retrospective, cohort-internal findings position DXA Agent as an inspectable, measurement-burden-aware research workflow requiring independent prospective validation.
Malec, S. A.; Pradhan, M.; Upadhayaya, R.; Metzger, V.
Show abstract
Objective: Observational studies are essential for investigating risk factors for Alzheimer's disease and related dementias (ADRD), but inconsistent reporting and selection of covariates can contribute to residual confounding, omitted-variable bias, and reduced reproducibility. We developed and evaluated VAREX (Variable Extraction), a large language model (LLM)-based information extraction framework designed to automatically identify exposures, outcomes, and covariates from epidemiologic studies and populate structured evidence repositories. Materials and Methods: VAREX combines retrieval-augmented generation, biomedical language-model embeddings, semantic chunking, cross-encoder reranking, and prompt-engineered LLM workflows to extract epidemiologic variables from full-text biomedical articles. The framework was evaluated using a reference-standard corpus of observational studies examining blood pressure variability (BPV) and Alzheimer's disease-related dementias (ADRD), together with external validation datasets involving other exposure-outcome relationships. Extracted variables were compared with independently curated human reference standards using semantic matching and one-to-one assignment procedures. Covariates were additionally classified into ten epidemiologically relevant semantic categories. Results: In the primary BPV[->]ADRD corpus (10 studies), VAREX achieved a precision of 0.91, recall of 0.84, and F1-score of 0.87 for variable extraction. Covariate classification accuracy was 0.90, yielding a strict extraction-and-classification F1-score of 0.78. External validation datasets demonstrated comparable performance across diverse epidemiologic domains, with extraction F1-scores ranging from 0.73 to 0.85. Category-level performance was strongest for health behaviors (F1=0.96), sociodemographic variables (F1=0.90), and medication exposures (F1=0.89). Compared with published estimates of manual systematic-review effort, VAREX reduced processing time from approximately 61 minutes to 9 minutes per article, representing an 85.7% reduction in review time. Discussion: These findings demonstrate that LLM-based information extraction can accurately identify and classify epidemiologic variables across heterogeneous observational-study designs. Automated extraction enables scalable construction of structured repositories of exposures, outcomes, and covariates while substantially reducing the labor required for evidence synthesis and systematic reviews. Conclusion: VAREX provides an effective framework for automated extraction and classification of epidemiologic variables from the biomedical literature. By supporting large-scale evidence synthesis and structured knowledge resource development, VAREX may facilitate more rigorous observational research, improved confounder identification, and enhanced reproducibility in epidemiology.
Cummins, J.; Drysdale, H.; Elson, M.; Hussey, I.; Goldacre, B.; DeVito, N. J.
Show abstract
Objective To evaluate the accuracy and cost of RegCheck, an automated large language model (LLM)-based workflow, for identifying clinical trial outcomes and detecting outcome misreporting by comparing its outputs with manual assessments from the COMPare Trials project. Design Validation study. Setting Sixty-two clinical trials originally assessed in the COMPare Trials project, sampled from five high impact general medical journals. Participants Published clinical trial reports and their corresponding prespecified registrations and/or protocols. Main outcome measures Four prespecified research questions were examined. RQ1 assessed outcome extraction recall relative to COMPare. RQ2 assessed accuracy of outcome classification as primary, secondary, or non-prespecified. RQ3 assessed accuracy of misreporting detection relative to COMPare, with additional manual adjudication of discrepancies between RegCheck and COMPare. RQ4 assessed the average per-paper cost of running the automated workflow. Results Across the validation papers, RegCheck achieved 91.2% outcome extraction recall relative to COMPare, and 83.6% outcome classification accuracy. For detection of outcome misreporting, RegCheck's overall accuracy was 85.6%. However, after resolving discrepancies with the original human judgements (which frequently favoured RegCheck's judgement), revised accuracy for outcome misreporting detection was 94.8%. The mean cost of running the workflow was 5.94 USD per paper. Conclusions RegCheck achieved high overall performance with a rigorous manual benchmark for identifying prespecified and reported outcomes in clinical trials, and detecting outcome misreporting, while operating at very low marginal cost. Adjudication of discrepant judgements suggested that RegCheck frequently identified valid issues not captured in the reference standard. Automated outcome checking may offer a scalable way to support editors, peer reviewers, and authors in detecting outcome switching and improving trial reporting.
Reddy, S.; Heritier, A.
Show abstract
The rapid expansion of the medical artificial intelligence (AI) literature has outpaced our ability to judge how far published models have progressed towards clinical use. We investigated whether the translational maturity of a study can be estimated automatically from its abstract. Using PubMed, we assembled a corpus of 11,024 candidate articles, reduced it to 1,816 AI-related articles by heuristic filtering, and manually double-annotated a balanced sample of 524 articles across five maturity classes (internal validation, external validation, prospective evaluation, implementation or governance, and not applicable). Abstracts were represented as TF-IDF features and classified using multinomial logistic regression with a Lasso penalty, chosen for interpretability and suitability for a small, imbalanced dataset. On a stratified held-out test set (n = 104), the model achieved 69.2% accuracy, Cohen's kappa of 0.495, macro-F1 of 0.458 and a weighted AUC of 0.820. Performance was strong for the frequent classes but poor for the rare implementation or governance class, which the model failed to recover. A balanced manual verification of 200 large-corpus predictions confirmed this pattern, with per-class precision ranging from 82.5% (internal validation) to 5.0% (implementation or governance). An interpretable, low-resource classifier can support literature mapping but requires human oversight for advanced maturity levels.
Park, C. S.-Y.
Show abstract
Background: Imputation accuracy is typically summarized by a single figure computed from one set of random deletions. For the heavy-tailed, collinear variables common in biomedical data, this study asks whether that figure can be trusted, and introduces a diagnostic protocol that determines when it cannot. Methods: This paper introduces GRIP (Geometry-aware Reproducibility of Imputation Protocol), a three-step diagnostic that profiles each variable's geometry and collinearity, stress-tests reproducibility under MCAR, MAR, and MNAR missingness with fixed seeds, and classifies the failure mode. GRIP was demonstrated on 1,885 United States Centers for Medicare & Medicaid Services (CMS) home-health agencies (15 numeric variables), comparing missForest, votingForest, mean, and median imputation. A supplementary k-nearest-neighbour (k-NN) grid (20 bivariate lognormal parameter combinations; n = 500) verified the SILENT failure boundary across imputer types. The simulation component followed the ADEMP framework. Results: Geometry profiling prospectively flagged two variables combining extreme right-skew (skewness 38.6, 43.3) with near-collinearity (|r| = 0.96). Under MCAR and MAR, missForest normalized root mean squared error (NRMSE) ranged from below 1 to 152 across otherwise identical replications (SD 14-30); the difficulty ordering of the two variables reversed in 69% of replications. Under MNAR self-masking, instability vanished (SD = 0), yet only 14% of true extreme magnitudes were recovered. Both failure modes arise from one mechanism: instability requires an extreme value to be absent while its collinear partner remains observed. A k-NN proxy grid confirmed SILENT failure in 11 of 15 high-skew parameter combinations under MNAR, regardless of correlation level. Conclusions: For heavy-tailed, collinear variables, one imputation-accuracy number can mislead in two opposite ways. GRIP detects both before an imputer is committed and is provided as reusable, open-source R code.
Ndabashinze, R.; Franzen, D.; Kozuch, E.; Aagerup, J.; Fink, A.; Yerunkar, S. S.; Hunter, K.; Mayo-Wilson, E.; Ying, X.; Kilicoglu, H.; Schorr, S. G.; Seidler, A. L.
Show abstract
Background Clinical trials conducted in Germany are registered across multiple registries, including the German Clinical Trials Register (DRKS), ClinicalTrials.gov, the EU Clinical Trials Register (EUCTR), and, since 2023, the Clinical Trials Information System (CTIS). These registries record health conditions using different classification systems and terminologies, including ICD-10-GM, MeSH, MedDRA, and free text, making cross-registry analyses difficult. We developed and evaluated a pipeline for harmonizing trial condition descriptions to WHO ICD-10 and compared its performance with that of a large language model (LLM) and to health conditions coded by humans. Methods We developed a four-stage, registry-aware mapping pipeline consisting of: (i) condition mention extraction and normalization; (ii) classification of ICD-mappable versus non-mappable mentions; (iii) ontology-based candidate generation using UMLS links between MeSH, MedDRA, ICD-10-GM, and WHO ICD-10; and (iv) SapBERT-based semantic retrieval with hybrid confidence scoring. A second variant additionally applied cross-encoder reranking of the top candidate codes. A stratified sample of 500 condition mentions was manually coded to create an expert reference standard. GPT-4o was evaluated in parallel using the same structured decision framework as the human reviewers. Performance was assessed using accuracy, precision, F1 score, and Cohen's {kappa} at the three-character, block, and chapter levels of ICD-10. Results The pipeline was applied to 23,061 clinical trials and identified 39,512 ICD-mappable condition mentions, of which 72.4% received a high-confidence assignment. Against 390 expert-coded mentions, the baseline pipeline achieved 49.0% accuracy at the three-character ICD-10 level ({kappa} = 0.487), increasing to 58.7% at the chapter level ({kappa} = 0.561). The cross-encoder method produced small but consistent improvements across all evaluation levels. Candidate-recall analysis showed that the correct code was present in the retrieved candidate set in only 73.7% of cases. The LLM substantially outperformed both pipeline variants, achieving 96.7% accuracy and near-perfect agreement with expert coding ({kappa} = 0.966) at the three-character level. The LLM also assigned clinically plausible codes to 82.4% of rejected mentions, 62.8% of Tier-3 exclusions, and 92.3% of review-band mentions. Conclusion Automated harmonization of clinical trial condition data across heterogeneous registries is feasible and supports the use of a common ICD-10 framework for cross-registry analyses. The LLMs achieved high agreement with expert coding, and performed better than the deterministic ontology and embedding pipeline, which achieved moderate agreement. These findings indicate that LLMs can support analyses of the distribution of health conditions investigated in clinical trials in Germany.They are a promising tool for classification of other non-standardised trial characteristics in registries. Keywords: Clinical trial registries; ICD-10; disease harmonization; UMLS; entity linking; SapBERT; large language models; clinical research; natural language processing.